Papers with MLE training

6 papers
Imitation Learning for Neural Morphological String Transduction (D18-1)

Copied to clipboard

Challenge: Recent studies have shown that neural transition-based models can be used for morphological tasks such as inflection generation and lemmatization without a character aligner or warm start.
Approach: They propose to use imitation learning to train a neural transition-based string transducer for morphological tasks such as inflection generation and lemmatization.
Outcome: The proposed model eliminates the need for a character aligner or warm start and achieves state-of-the-art performance on several datasets.
Energy-Based Reranking: Improving Neural Machine Translation Using Energy-Based Models (2021.acl-long)

Copied to clipboard

Challenge: Autoregressive neural machine translation (NMT) uses a tractable likelihood computation and efficient sampling.
Approach: They propose to use an energy-based model to mimic the behavior of the task measure and use it to train an energy based re-ranking algorithm.
Outcome: The proposed model improves on the samples drawn from the NMT with a higher BLEU score than the experimental model and the energy-based re-ranking algorithm.
Helping the Weak Makes You Strong: Simple Multi-Task Learning Improves Non-Autoregressive Translators (2022.emnlp-main)

Copied to clipboard

Challenge: Non-autoregressive (NAR) neural machine translation models require a conditional independence assumption on target sequences, resulting in less informative learning signals.
Approach: They propose a model-agnostic multi-task learning framework to provide more informative learning signals for NAR models under conventional MLE training.
Outcome: The proposed framework improves accuracy of multiple NAR baselines without additional decoding overhead.
CaLcs: Continuously Approximating Longest Common Subsequence for Sequence Level Optimization (D18-1)

Copied to clipboard

Challenge: Maximum-likelihood estimation (MLE) is widely used for text-generation based natural language processing applications.
Approach: They propose a method to train models with maximum-likelihood estimation using a differentiable surrogate of longest common subsequence measure that captures sequence-level structure similarity.
Outcome: Experimental results show that the proposed approach improves on the current MLE approach for downstream tasks like text summarization and machine translation.
Diverse Keyphrase Generation with Neural Unlikelihood Training (2020.coling-main)

Copied to clipboard

Challenge: Recent advances in neural natural language generation have made possible remarkable progress on the task of keyphrase generation, however, the importance of diversity in keyphrases has been largely ignored.
Approach: They propose to train a sequence-to-sequence keyphrase generation model from the perspective of diversity.
Outcome: The proposed model achieves large diversity gains while maintaining competitive output quality.
Less Likely Brainstorming: Using Language Models to Generate Alternative Hypotheses (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods to reduce cognitive errors in MRI interpretations do not work for generating less likely outputs.
Approach: They propose a task that asks a model to generate outputs that humans think are relevant but less likely to happen.
Outcome: The proposed method compares with several state-of-the-art controlled text generation models via automatic and human evaluations and shows that it reduces cognitive errors in interpreting MRI findings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations